Goto

Collaborating Authors

 tennis court


Learning Wheelchair Tennis Navigation from Broadcast Videos with Domain Knowledge Transfer and Diffusion Motion Planning

arXiv.org Artificial Intelligence

In this paper, we propose a novel and generalizable zero-shot knowledge transfer framework that distills expert sports navigation strategies from web videos into robotic systems with adversarial constraints and out-of-distribution image trajectories. Our pipeline enables diffusion-based imitation learning by reconstructing the full 3D task space from multiple partial views, warping it into 2D image space, closing the planning loop within this 2D space, and transfer constrained motion of interest back to task space. Additionally, we demonstrate that the learned policy can serve as a local planner in conjunction with position control. We apply this framework in the wheelchair tennis navigation problem to guide the wheelchair into the ball-hitting region. Our pipeline achieves a navigation success rate of 97.67% in reaching real-world recorded tennis ball trajectories with a physical robot wheelchair, and achieve a success rate of 68.49% in a real-world, real-time experiment on a full-sized tennis court.


Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models

arXiv.org Artificial Intelligence

Object hallucination has been an Achilles' heel which hinders the broader applications of large vision-language models (LVLMs). Object hallucination refers to the phenomenon that the LVLMs claim non-existent objects in the image. To mitigate the object hallucinations, instruction tuning and external model-based detection methods have been proposed, which either require large-scare computational resources or depend on the detection result of external models. However, there remains an under-explored field to utilize the LVLM itself to alleviate object hallucinations. In this work, we adopt the intuition that the LVLM tends to respond logically consistently for existent objects but inconsistently for hallucinated objects. Therefore, we propose a Logical Closed Loop-based framework for Object Hallucination Detection and Mitigation, namely LogicCheckGPT. In specific, we devise logical consistency probing to raise questions with logical correlations, inquiring about attributes from objects and vice versa. Whether their responses can form a logical closed loop serves as an indicator of object hallucination. As a plug-and-play method, it can be seamlessly applied to all existing LVLMs. Comprehensive experiments conducted on three benchmarks across four LVLMs have demonstrated significant improvements brought by our method, indicating its effectiveness and generality.


Efficient Remote Sensing with Harmonized Transfer Learning and Modality Alignment

arXiv.org Artificial Intelligence

With the rise of Visual and Language Pretraining (VLP), an increasing number of downstream tasks are adopting the paradigm of pretraining followed by fine-tuning. Although this paradigm has demonstrated potential in various multimodal downstream tasks, its implementation in the remote sensing domain encounters some obstacles. Specifically, the tendency for same-modality embeddings to cluster together impedes efficient transfer learning. To tackle this issue, we review the aim of multimodal transfer learning for downstream tasks from a unified perspective, and rethink the optimization process based on three distinct objectives. We propose "Harmonized Transfer Learning and Modality Alignment (HarMA)", a method that simultaneously satisfies task constraints, modality alignment, and single-modality uniform alignment, while minimizing training overhead through parameter-efficient fine-tuning. Remarkably, without the need for external data for training, HarMA achieves state-of-the-art performance in two popular multimodal retrieval tasks in the field of remote sensing. Our experiments reveal that HarMA achieves competitive and even superior performance to fully fine-tuned models with only minimal adjustable parameters. Due to its simplicity, HarMA can be integrated into almost all existing multimodal pretraining models. We hope this method can facilitate the efficient application of large models to a wide range of downstream tasks while significantly reducing the resource consumption. Code is available at https://github.com/seekerhuang/HarMA.


A Appendix

Neural Information Processing Systems

A.1 Compute Usage The seven billion parameter language model we used as part of Frozen used model parallelism with the strategy from [39] to partition one instance of the model over four accelerators. Each instance had a batch size of 8. To reach a batch size of 128 in this configuration, we additionally employed data parallelism with 16 synchronous replicas. The whole system was trained on a 4x8 TPUv3 [15] topology for about 12 hours, which is when validation set performance for Conceptual Captions led us to do early stopping. A.2 Frozen Architecture Details The pretrained transformer language model we used has a GPT-like architecture [30]. It consists of a series of identical residual layers, each comprised of a self-attention operation followed by a positionwise MLP.


Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

arXiv.org Artificial Intelligence

Large vision-language models (LVLMs) have shown remarkable abilities in understanding visual information with human languages. However, LVLMs still suffer from object hallucination, which is the problem of generating descriptions that include objects that do not actually exist in the images. This can negatively impact many vision-language tasks, such as visual summarization and reasoning. To address this issue, we propose a simple yet powerful algorithm, LVLM Hallucination Revisor (LURE), to post-hoc rectify object hallucination in LVLMs by reconstructing less hallucinatory descriptions. LURE is grounded in a rigorous statistical analysis of the key factors underlying object hallucination, including co-occurrence (the frequent appearance of certain objects alongside others in images), uncertainty (objects with higher uncertainty during LVLM decoding), and object position (hallucination often appears in the later part of the generated text). LURE can also be seamlessly integrated with any LVLMs. We evaluate LURE on six open-source LVLMs, achieving a 23% improvement in general object hallucination evaluation metrics over the previous best approach. In both GPT and human evaluations, LURE consistently ranks at the top. Our data and code are available at https://github.com/YiyangZhou/LURE.


Improving Image Clustering through Sample Ranking and Its Application to remote--sensing images

arXiv.org Artificial Intelligence

Image clustering is a very useful technique that is widely applied to various areas, including remote sensing. Recently, visual representations by self-supervised learning have greatly improved the performance of image clustering. To further improve the well-trained clustering models, this paper proposes a novel method by first ranking samples within each cluster based on the confidence in their belonging to the current cluster and then using the ranking to formulate a weighted cross-entropy loss to train the model. For ranking the samples, we developed a method for computing the likelihood of samples belonging to the current clusters based on whether they are situated in densely populated neighborhoods, while for training the model, we give a strategy for weighting the ranked samples. We present extensive experimental results that demonstrate that the new technique can be used to improve the State-of-the-Art image clustering models, achieving accuracy performance gains ranging from $2.1\%$ to $15.9\%$. Performing our method on a variety of datasets from remote sensing, we show that our method can be effectively applied to remote--sensing images.


Mind blown by AI art? Wait to you see AI-generated video

#artificialintelligence

AI art has been bursting into the mainstream thanks to the likes of DALL-E 2 and MidJourney. The tools allow anyone to create almost any image they can dream of from just a short text prompt. The results can be very, very strange, but artists, designers and brands are learning how to make the technology work for them, sometimes very successfully. But if you've been impressed so far, it seems the next advance is already on the way: AI video generators. AI art generators work based on text prompts. You type in what you want, and the art generator will create it โ€“ or at least its interpretation of the prompt.


Is it time for cutting-edge tech to make your mower greener?

The Guardian

Gardeners want to make their grass even greener. As petrol prices rocket and people become ever more conscious of their environmental impact, many are turning to the latest generation of lawnmowers to keep their gardens looking good. While the fronts of our houses are gradually seeing the replacement of petrol cars with electric vehicles, advances in lithium-ion batteries have meant that the trusted back garden mower has also been given a modern overhaul โ€“ but at a price. So is it time to replace your current mower with a battery-powered or "robot" version, stick with petrol despite the spiralling costs, or stay plugged in? The length and breadth of your garden will heavily influence what type of machine you need.


Motion-capture software used in Hollywood helps tennis athletes stay healthy

Daily Mail - Science & tech

Serena Williams, Roger Federer, Rafael Nadal, Novak Djokovic and Andy Murray have all struggled with serious injury. But future tennis stars could potentially be spared such discomfort with the help of new motion-capture technology, according to bio-mechanical experts from Coventry University. The software uses 3D optical tracking equipment, similar to that used for Hollywood movies, and then applies its own algorithms to measure loads imposed on joints, bones and muscles, according to the team. The resulting data can help improve training and, crucially, it could help players avoid injury. With Wimbledon getting into full swing next week, researchers have developed a motion capture system that can help tennis players avoid injury.


The US Again Has World's Most Powerful Supercomputer

WIRED

Plenty of people around the world got new gadgets Friday, but one in Eastern Tennessee stands out. Summit, a new supercomputer unveiled at Oak Ridge National Lab is, unofficially for now, the most powerful calculating machine on the planet. It was designed in part to scale up the artificial intelligence techniques that power some of the recent tricks in your smartphone. America hasn't possessed the world's most powerful supercomputer since June 2013, when a Chinese machine first claimed the title. Summit is expected to end that run when the official ranking of supercomputers, from an organization called Top500, is updated later this month.